Repository navigation
Expose bundle-aware speculative defaults before model load - #3029
Conversation
Points the app at vmlx-swift 371f5c40: multi-row bit-exact BF16-affine verify kernels, lane matmul with safe tiling and the DFlash2 width chooser, native-MTP copy drafts, and the Qwen4 PLE page cache. Proof build only; not for release.
Pins vmlx-swift perf/claude-swift-port-oct6 @ a4f99a73 (six pin sites).
Settings > Speculative Decoding now has three modes: Off (AR), Default
and On (Adaptive). Default is the engine's `.familyDefault`: Adaptive
native MTP for Qwen3.8 Flash-Next, and a Qwen 27B bundle's own dflash2/
drafter. A fresh install shows the chat picker's Native MTP row as On
for Flash-Next bundles, and the user can switch it off. The picker and
the phone snapshot show the per-bundle effective state.
ModelRuntime passes the bundle directory to drafter selection, so a
bundled dflash2/ drafter is found without a folder pick.
The Safe Auto materialized-load memory check is advisory: it logs a
warning instead of refusing. It refused Flash-Next JANG_4S on a 128 GB
Mac ("require ~92 GiB, only ~53 GiB available"); the model loads and
runs at 55-115 tok/s. Strict mode keeps its explicit refusal.
Pins vmlx-swift perf/claude-swift-port-oct6 @ 0e26bcb5 (six pin sites). The engine now loads dense Qwen3.5 JANGH bundles (Qwen3.8-27B-JANGH2: jangtq2 2-bit MLP banks). Measured in RunBench on max2: AR 30.6 tok/s prose, DFlash2 28.6 / 57.7 / 135.0 / 141.0 tok/s (prose / code / easy code / easy prose).
|
Current live qualification found a numerical blocker; this PR remains draft and unmerged. Source tested: Osaurus ffb03f2, vmlx-swift fd9d925c, MLX fork5aba1efd. The later engine91726781 change is whitespace only; it is not claimed as a new app build. Live isolated Release app, dense Qwen3.8-27B-JANGH2:
Controlled app HTTP diagnostic, same38-token prose, explicit greedy, fresh cache namespace, AR/Adaptive/Adaptive/AR:29.01 /29.39 /27.37 /26.44 tok/s. Clocks differed, so this does not establish a speedup. AR answers matched each other; Adaptive answers differed from AR and from each other. Holding DFlash width at its trained8 produced identical outputs twice at27.61/27.46 tok/s, but still differed from AR. This is a diagnostic, not a proposed fixed-depth product change. Causal teacher-forced test on the actual bundle: from the same prefilled snapshot, target row0 of widths5/8/16 differs from a one-token call before acceptance or rollback, both with and without LaneQMM and in eager/staged modes (12 comparisons, max absolute logit difference0.125). Lane installation also changes cold-prefill logits. Next work isolates the first differing operator and tests an arithmetic repair. Changing the controller alone is insufficient. Local raw artifacts: /Users/eric/vmlx-private-evidence/qwen-spec-review-20261007/{app-live002.log,ui-dense-prose.png,ui-dense-followup.png,dense-prose-abba001,dense-fixed8-001,review/prefill-partition/dense001.json,review/DFLASH-DENSE-ROW-CAUSAL.md}. These local artifacts are not publicly accessible CI attachments. Other quants, complete media/cache/lifecycle coverage, and final app proof remain pending. |
|
Correction to the preceding qualification comment: I applied the wrong numerical contract to dense27B DFlash2. The supplied Swift handoff READ-FIRST §5c explicitly documents NAX/Lane verification as speed-first and not bit-identical to AR. The strict AR row-exact gate applies to Flash-Next native MTP. The measured S1-versus-wide differences are real, but they are not evidence of a new regression or a standalone merge blocker under that documented dense27B contract. No production kernel or adaptive-width change was made. The fixed8 run was diagnostic only. Review resumes against the correct references: fast QMV vs its replaced S1 kernel; small NAX tile vs its replaced large tile; verification/acceptance and cache commit within the chosen execution path; default selection, live app speeds, cold prefill, tool continuation, and media fallback. Dense27B controlled prose27–29 tok/s is consistent with the handoff28.6 tok/s. The23 tok/s visible-composer run had4117 prompt tokens and is not a matched regression result. The PR remains unmerged pending the remaining integration/CI proofs, not because dense DFlash differs from plain AR. |
A working-set estimate that meets the Safe Auto budget left a 0-byte allocator headroom, so native-MTP and DFlash 2 generations ran with no freed-buffer reuse: every decode step re-allocated its intermediates and paid a Metal residency commit per buffer. The working set already prices a >= 2 GiB scratch floor for exactly these buffers, so the remaining-budget clamp may shrink the pool but never below that floor. Measured live on Qwen3.8 Flash-Next JANG_4S under Safe Auto: 31.4 ms per target forward with the starved pool vs 21.6 ms with a working pool.
…e, DFlash 2 AR width
Live app proof — osaurus dev build pinned to this engine headSetup.
AR references on the same machine: 27B 4D ~23, 27B JANGH2 ~29, Flash-Next ~50–55 tok/s. "Cold" = the first request App-side defect found and fixed in osaurus (
|
Pin → vmlx-swift
|
The family default loads without the native MTP head for every family except Qwen3.8 Flash-Next. Switching a resident model from the family default to On (Adaptive) compared only off/not-off, so no reload ran and the head-less model never speculated until a manual reload. A change into or out of the family default is now a load-input change (ServerControllerConfigLoadingTests already expected this and failed in CI). Tests brought up to date with the shipped engine contracts: - the settings default is the bundle-aware family default, not Off; - a selectable external drafter needs complete weights (engine validates safetensors headers against the required shapes), so the fixture writes a sparse but complete artifact instead of a config-only folder; - the picker blocked-bundle check follows the renamed local in FloatingInputCard (behaviour unchanged: blocked bundles render Off).
|
Ready for final audit at |
|
Independent audit update at app Live app proof on Qwen3.8-27B-JANGH2:
Remaining: speed comparison is not cleared. Current sampled full-matrix JANGH2 medians are35.36 prose/93.29 code/149.91 easy-code tok/s, below parts of the historical table; additional runs show clock and adaptive-width variation. No sampler/kernel changes have been made to chase these numbers. Isolated profile also exposed history-save errors in code unchanged from main, and duplicate bundle basenames make the short Evidence receipts: |
|
Additional independent live app audit on f41de0a / engine e54e7388:
Found one misleading baseline diagnostic: nil nativeMTPStats logs plain/off even for actual DFlash2 cycles. A diagnostics-only correction is under build verification; no decoder/kernel/cache/sampler changes. Merge remains held: current H2 sustained prose speed is not cleared against historical receipts, and remaining quant coverage is incomplete. Raw private receipts: pr565-pr3029-audit-20261007/app-live-restarted.log, app-media-prefill-debug.log, media-cache-stats.json, media-physical-footprint.txt. |
|
Diagnostics correction pushed as 9167afe: nil native-MTP stats no longer claim plain AR. The line now states nativeMTP=off decodePath=unreported, leaving actual DFlash path proof to engine telemetry. No generation, cache or kernel behavior changed. Fresh isolated Release build passed. In the rebuilt app, attached a neutral-named image with yellow2/blue background after the previous white7/red image. Both the initial answer and follow-up were correct and completed naturally. Changed media used a different salt; follow-up restored2340 tokens with111 suffix tokens and aligned drafter offsets. New diagnostic text is observed live alongside real DFlash cycles. App history survived relaunch. These short correctness rows overlapped CPU compilation and are not speed gates:15tokens22.6tok/s and8tokens13.9tok/s. Remaining sustained-speed audit and quant coverage are still open; not merge-ready yet. Receipts: app-telemetry-build.log, app-telemetry-live.log, app-telemetry-prefill-debug.log in private pr565-pr3029-audit-20261007 evidence. |
|
Independent current-engine checks (still partial; not a merge-ready claim):
SOURCE EVIDENCE: engine resolver/current headers, app bundle-capability picker path; bundle-resolution-receipt.json records binarySHA and exact paths. |
|
Pushed 129 focused tests passed across NativeMTPPreloadDetectionTests, NativeMTPAdmissionTests and RuntimePolicySourceTests before the mechanical engine pin update. Added adversarial Qwen dense/Flash affine/JANGH aliases, malformed/missing config, conflicting nested architecture, and explicit Off assertions. All six tracked runtime pin sites now point to engine SOURCE EVIDENCE: ModelFamilyNames.swift, ModelRuntime.swift and NativeMTPPreloadDetectionTests.swift at |
The subagent runner always set prompt_tokens from the message estimator, which omits the rendered chat template and the separately attached tool schemas. A 4-tool delegated child reported 92 prompt tokens while the runtime prefilled 1,455 (same 1,427 completion tokens / 48.5 tok/s), and total, worker and context-saved figures inherited the gap. The runner now takes the count from the input-token hint the runtime already streams (or the stats sentinel's input count) and keeps the estimator only as the fallback when neither arrives. Completion and tokens-per-second are unchanged. Tests: runtime count wins over the estimate; estimator fallback when no count is streamed.
Audit follow-up — head
|
|
Final consuming audit: app 15e917f pins merged engine a468573fbe5b2c166d8d1dabea78c205c41e6716 at all six tracked sites. Engine #565 is merged. The merge tree exactly matches tested engine f3963efb7. SOURCE EVIDENCE: the final app delta is only those six pin replacements; all existing performance, mode-reload, drafter discovery, allocator scratch-floor and delegation usage changes remain in this PR. The 30-second inactivity setting/toggle is unchanged. Engine final audit adds only split-fusion completeness and bounds-safe residency metadata accounting; it does not change decode kernels, sampling, cache storage or adaptive scheduling. LIVE EVIDENCE / TESTS:
Raw current artifacts: private merge-final-delta-20261007/app-build.log, app-tests.log, real-bundles-{family_default,off}.log and receipt. Frozen prior live evidence is indexed under pr565-pr3029-audit-20261007/claude-live-consolidation/README.md. The bounded earlier 27B H2 Off delegation timeout is not represented as a completed-answer pass. Restored TTFT is not cold-prefill throughput. GitHub CI is separately reported by its checks: running/queued checks are not claimed green. Local affected tests above passed at the exact consuming pin; engine repository-wide formatter drift and inherited quant-metadata tests remain documented. No release, tag, appcast, model-unload policy change or new performance tuning is included. |
|
Merged after all app CI checks passed: https://github.com/osaurus-ai/osaurus/actions/runs/37704672816. Osaurus main is Local proof: 258 affected app tests, 6 engine delta tests, 9 real bundle Default/Off resolver pairs. Existing live speed/cache/media/delegation evidence is retained with the limits in the preceding audit comments. 30-second inactivity setting unchanged. No release/tag/appcast. Internal final handoff and hashed stop receipt are saved; work stopped. |
Selecting a supported Qwen bundle exposes speculative On (Adaptive) / Off (AR) before the model loads.
Engine pin: merged osaurus-ai/vmlx-swift#565 at
a468573fbe5b2c166d8d1dabea78c205c41e6716, updated at all six tracked sites. This includes the performance work, follow-up drafter cache repair, physical head validation, SSD n-gram defaults, residency accounting, and the final split-fusion/invalid-offset audit fixes.App-side changes in this PR
492c66f8). Under Safe Auto the remaining-budget clamp could set MLX's freed-buffer pool to 0 for a large model, so every decode step re-allocated with a residency commit. Flash-Next 4S target forward went from 31.4 to 20.5 ms, live.f41de0a3). The family default loads without the MTP head for every family except Flash-Next, so moving between the family default and On (Adaptive) is now a load-input change. Before, a resident model could stay head-less and never speculate.Evidence
HTTPHandlerEndpointTests.runtimeSettings_put_persistsAndReportsRuntimeEffectscan fail when run in parallel with other suites that override the shared settings directory. It passes alone (3/3) and in CI.Known behaviour, not a defect
Each reasoning on/off/effort setting keeps its own prefix-cache chain (it is part of the cache key), so the first turn after a change re-prefills.
No release is included. Final consuming-build audit evidence is recorded in the merge comments.